Papers with Restrictive Setting
Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Large Language Models' (LLMs) performance is greatly inflated by memorization, a study finds . authors propose a framework to analyze how LLMs reason under different degrees of memory access. |
| Approach: | They propose a framework that allows large language models to reason under different degrees of memory access. |
| Outcome: | Evaluating GPT-4o, LLaMA3.3-70B, and DeepSeek V3 on character-centric story understanding benchmarks, they find up to a 45.2% accuracy drop under the Restrictive Setting. |